Papers with convergence process
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Jailbreak Attacks without Compromising Usability (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for enhancing LLM security compromise usability, study finds . boundary-safe representations close to harmful representations are disrupted, resulting in usability decline . |
| Approach: | They propose a method to push harmful representations away from boundary-safe representations and obtain an exact distinction boundary. |
| Outcome: | The proposed method reduces over-refusal rate and maintains general capability . it pushes harmful representations away from boundary-safe representations, thereby reducing usability. |
Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that large language models are struggling to align with human preference in complex tasks and scenarios. |
| Approach: | They propose a low-redundant alignment method that selects the top-10% most updated parameters in LLMs for alignment training. |
| Outcome: | The proposed method improves on 10 datasets and shows that it is redundant . it can be used to train LLMs on QA and ECQA datasets, but it is not feasible to test it on a large dataset. |